Papers by Vu Minh Hieu Phan
MMCLIP: Cross-Modal Attention Masked Modelling for Medical Language-Image Pre-Training (2026.acl-long)
Copied to clipboard
| Challenge: | Existing vision-and-language pretraining methods face challenges in reconstructing pathological features due to limited data. |
| Approach: | They propose a method that uses masked modeling to enhance visual and linguistic learning. |
| Outcome: | MMCLIP integrates unpaired data through disease-kind prompts to achieve state-of-the-art performance in zero-shot and fine-tuning across five benchmarks. |